Back

Medical Decision Making

SAGE Publications

Preprints posted in the last 7 days, ranked by how well they match Medical Decision Making's content profile, based on 12 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Primary Care Quality and Inappropriate Community Antibiotic Use: A Double Machine Learning Instrumental Variable Approach

Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.

2026-08-31 health economics 10.64898/2026.08.26.26361459 medRxiv
Top 0.1%
5.1%
Show abstract

Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.

2
Increasing Lung Cancer Screening Participation Using an Informational Video Nudge: A Randomized Feasibility Trial

Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.

2026-09-01 health systems and quality improvement 10.64898/2026.08.28.26361654 medRxiv
Top 0.1%
1.8%
Show abstract

Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.

3
Modeling the health and economic impact of scaling up monthly oral pre-exposure prophylaxis alone or alongside injectable Lenacapavir in Kenya and South Africa

Malhotra, A.; Patel, N.; Kaftan, D.; Mudimu, E.; Bershteyn, A.; Sharma, M.

2026-09-04 health economics 10.64898/2026.09.01.26361983 medRxiv
Top 0.2%
1.4%
Show abstract

Introduction: Monthly oral HIV pre-exposure prophylaxis (PrEP) such as MK-8527 offer promise as low-cost, self-administered options that are easily delivered through community-based platforms. Economic evaluations of MK-8527 either alone or alongside other long-acting (LA) PrEP products like lenacapavir are needed for informing HIV prevention strategies. Methods: We adapted an agent-based network model, EMOD-HIV, to simulate LA-PrEP scale-up in South Africa and western Kenya from 2026-2035; scenarios evaluated MK-8527 alone, lenacapavir alone and combined strategies, with varying uptake among female sex workers, their male clients, and individuals with >1 partner. We assumed 95% effectiveness of MK-8527 for 2 months (assuming individuals took 2 of 3 pills dispensed) and 95% lenacapavir effectiveness for 6 months. Scenarios were compared to a baseline of daily oral PrEP only. Results: Assuming the same uptake rates, MK-8527 alone achieved lower health impacts than lenacapavir alone in both settings (6-14% vs. 11-18% of HIV infections averted) but had substantially lower costs; provision costs of MK-8527 were 58-59% lower than lenacapavir assuming $1.00/pill and 40-43% lower at $2.50/pill. In western Kenya, ICERs for MK-8527 alone were $467/DALY averted and $799/DALY averted assuming pill prices of $1.00 and $2.50 respectively, compared to $1,306/DALY averted for lenacapavir. In South Africa, all LA-PrEP strategies were cost-saving over the 35-year horizon, although near-term budget impacts were substantial ($169-348 million over five years). Service delivery accounted for the majority of MK-8527 costs (73% at $1.00 per pill). Combined strategies of lenacapavir and MK-8527 increased health benefits (13-23% infections averted) but had higher provision costs than either strategy alone. Conclusion: MK-8527 can reduce HIV incidence at lower costs than lenacapavir. However, high service delivery costs limit its cost-effectiveness to scenarios in which pill prices are low (US$1.00 per pill) and provision is targeted to populations at substantial HIV risk.

4
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.2%
1.2%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

5
New tests for trials of very few patients using longitudinal data - a case-study in Autosomal Recessive Cerebellar Ataxias

Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.

2026-09-02 health informatics 10.64898/2026.08.28.26361588 medRxiv
Top 0.3%
1.1%
Show abstract

We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.

6
Individual-Level Counterfactual Analysis of SGLT2 Inhibitors Versus DPP4 Inhibitors in Diabetic Kidney Disease Using Causal Machine Learning

Yano, Y.; Nagasu, H.; Hiroshi, K.; Ohashi, M.; Isaka, Y.; Okada, H.; Nangaku, M.; Kashihara, N.

2026-09-03 health informatics 10.64898/2026.08.30.26361750 medRxiv
Top 0.3%
1.1%
Show abstract

Background: Traditional real-world studies comparing SGLT2 and DPP4 inhibitors on renal outcomes rely on propensity score matching, which causes high-dimensional data loss. We used causal machine learning (Causal ML) to unmask heterogeneous treatment effects in diabetic kidney disease (DKD). Methods: Using data from 4,588 patients within the Japanese J-CKD-DB-Ex registry, we implemented a doubly robust (DR) learning framework (Linear DR-learner with XGBoost) to compare SGLT2 and DPP4 inhibitors. Outcomes included the chronic eGFR slope and a composite renal endpoint ([≥] 50% eGFR decline or end-stage kidney disease). Heterogeneity was explored via causal SHAP and decision trees. Results: At the population level, SGLT2 inhibitors modestly slowed chronic eGFR decline (average treatment effect [ATE] = 0.14 [95% CI: -0.86, 1.15] mL/min/1.73m^2/year) and reduced composite endpoint risk by 9% (ATE: -0.09 [-0.11, -0.08]) versus DPP4 inhibitors. However, individual-level counterfactual analysis suggested that for the chronic eGFR slope, non-glinide users with stable pre-treatment trajectories who were also taking ACE inhibitors had a greater benefit from SGLT2 inhibitors (ATE: 2.95 [-0.68, 6.58]). Conversely, glinide users with steep pre-treatment decline had a greater benefit from DPP4 inhibitors (ATE: -8.98 [-16.11, -1.85]). For composite renal events, SGLT2 inhibitors had a 28% absolute risk reduction within the algorithmically identified high-risk subgroup (eGFR [≤] 28.1 mL/min/1.73 m^2 and positive proteinuria; ATE: -0.28 [-0.33, -0.23]). Even non-proteinuric decliners demonstrated a 8% risk reduction with SGLT2 inhibitors (ATE: -0.08 [-0.10, -0.06]). Conclusion: Causal ML advances precision medicine in DKD, shifting from uniform prescribing to individualized, data-driven therapy targeting distinct intrarenal pathways.

7
Effects of collaborative clinical visit agenda-setting interventions: A systematic review and meta-analysis

Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.

2026-09-03 medical education 10.64898/2026.08.30.26361729 medRxiv
Top 0.4%
0.6%
Show abstract

Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.

8
Artificial Scientific Intelligence for Measurement-burden-aware Modelling and Interpretation of Multi-site Bone Mineral Density

Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.

2026-09-01 health informatics 10.64898/2026.08.30.26361665 medRxiv
Top 0.7%
0.4%
Show abstract

Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.

9
A Multi-Agent Large Language Model Reasoning Engine for Early Detection of Pediatric Growth Disorders

Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.

2026-08-31 health informatics 10.64898/2026.08.28.26361655 medRxiv
Top 0.7%
0.4%
Show abstract

Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.

10
Multi-season evaluation and analysis of categorical trend forecasts of influenza hospital admissions in the United States

Davis, J. T.; Kaur, G.; Hines, A.; Ben-Nun, M.; Venkatramanan, S.; Brooks, L.; Mathis, S.; Ajelli, M.; Litvinova, M.; Kummer, A. G.; Ventura, P. C.; Mhade, S.; Weber, D.; Shemetov, D.; DeFries, N.; McDonald, D. J.; Yamana, T.; Zepeda-Tello, R.; Shaman, J.; Yaari, R.; Pei, S.; Webber, A.; Shandross, L.; Ray, E.; Wadsworth, S.; Niemi, J.; Redman, W. T.; Mullany, L.; Posner, R.; Mallela, A.; Lin, Y. T.; Hlavacek, W. S.; Smart, A.; Gill, A. A.; Drennan, A.; Fiebiger, B. J.; Miller, E. F.; Lee, J.; Mihaljevic, J. R.; Geist, K. A.; Baltz, M.; Bernik, O.; Truong, Y.-M. B.; Chen, Y.; Grosvenor, C. J.;

2026-09-02 epidemiology 10.64898/2026.08.31.26361843 medRxiv
Top 0.7%
0.4%
Show abstract

Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making.

11
Cost-Utility Analysis of First-Line Olaparib plus Abiraterone for Metastatic Castration-Resistant Prostate Cancer in China after Volume-Based Procurement

SHI, J.; Gu, Q.; Pan, J.; Yang, A.; Fan, M.

2026-08-31 health economics 10.64898/2026.08.26.26361314 medRxiv
Top 0.8%
0.3%
Show abstract

To evaluate the cost-utility and 5-year budget impact of first-line olaparib plus abiraterone versus abiraterone alone for metastatic castration-resistant prostate cancer (mCRPC) in China after the eleventh round of volume-based procurement (VBP). The intention-to-treat (ITT) population was assigned primary decision-analytic weight; the prespecified BRCA1/2-mutated (BRCAm) subgroup was a supporting analysis.

12
Multidimensional Social Vulnerability and Hepatic and Extrahepatic Outcomes in Adults With HIV/HBV Coinfection in the United States

Yendewa, G.; Chengsupanimit, T.; Dehghani, A.; Ahmed, A.; Mohareb, A.; Cohen, C.; Freeman, M.; Kim, H. N.; Ofotokun, I.; Dube, K.

2026-09-02 hiv aids 10.64898/2026.08.31.26361853 medRxiv
Top 0.8%
0.3%
Show abstract

Background: HIV/HBV coinfection is associated with substantial liver-related morbidity and mortality, yet the impact of social vulnerability (SV) on clinical outcomes has not been systematically assessed. We evaluated associations of multidimensional SV with mortality, hepatic, virologic, and extrahepatic organ outcomes among adults with HIV/HBV. Methods: We conducted a retrospective cohort study using TriNetX data from 110 U.S. healthcare organizations (2010-2026). We propensity score matched adults with HIV/HBV with and without documented SV 1:1 (2,024 per group). SV was defined using a four-domain framework encompassing material, healthcare access and engagement, interpersonal, and psychosocial vulnerability. Results: Over 15,900 person-years, SV was associated with higher mortality (hazard ratio [HR], 2.06; 95% confidence interval [CI], 1.72-2.47), liver composite events (HR, 1.37; 95% CI, 1.07-1.76), hepatic decompensation (HR, 1.94; 95% CI, 1.39-2.70), hepatic failure (HR, 2.39; 95% CI, 1.53-3.73), HBV viremia (HR, 1.69; 95% CI, 1.32-2.16), and HIV viremia (HR, 2.05; 95% CI, 1.71-2.46). SV was also associated with major adverse cardiovascular events (HR, 1.47), chronic kidney disease (HR, 1.49), and diabetes (HR, 1.25). Multidomain SV generally showed stronger associations than single-domain SV for most hepatic and virologic outcomes, with HR ranges of 1.76-2.62 versus 1.35-1.76 for single-domain SV. Healthcare access and engagement vulnerability was most consistently associated with mortality and hepatic outcomes. Conclusions: SV was associated with mortality, hepatic disease, impaired HIV/HBV control, extrahepatic organ morbidity, and acute care utilization in adults with HIV/HBV. SV assessment may improve risk stratification and identify actionable intervention targets during HIV/HBV care.

13
Cost-Aware Active Feature Acquisition for Differential Diagnosis under Realistic Clinical Availability Constraints

Bingham, J. C.; Arussy, N.

2026-08-31 health informatics 10.64898/2026.08.30.26361745 medRxiv
Top 0.8%
0.3%
Show abstract

Active Feature Acquisition (AFA) adaptively selects which diagnostic test to order next and offers a route to reduce unnecessary laboratory testing in acute care. Existing clinical AFA evaluations, however, assume every feature can be retrieved on demand and split data at the visit level, both of which inflate apparent performance. We re-evaluate cost-aware AFA under constraints designed to reflect deployment. From MIMIC-IV we constructed a cohort of 64,766 acute admissions (39,884 patients; 21 conditions; 55 features in 30 test panels) with a patient-level split, a 12-hour decision cutoff, and a per-patient availability mask from what was actually measured, and priced panels using the 2026 Medicare fee schedule under panel-level billing. We evaluated EIG-Cost, which scores each panel by Monte-Carlo Expected Information Gain penalised by its dollar cost, against eight published methods across budgets \30--$60 over five patient-level resamples. At a $30 budget, EIG-Cost achieved the highest macro-F1 (0.188, 95% CI [0.185, 0.191]) at the lowest cost ($17.28), exceeding the strongest baseline in all five resamples (p<0.001; Cohen's d=4.0), and led at every budget. Three of the eight methods collapsed to a vitals-only baseline (macro-F1 approx 0.040), acquiring nothing even at higher budgets, a genuine failure to adapt to availability rather than a budget limitation. Despite modest absolute accuracy, EIG-Cost's probabilities were well-calibrated (expected calibration error $0.048$). Under realistic availability constraints, clinical AFA is substantially harder than full-availability benchmarks imply, several published methods fail outright, and cost-aware information-gain scoring is a robust choice in this harder setting.

14
Limits of Single-Pass Retrieval-Augmented Generation for AI-Powered Cancer Care Navigation: A Comparison of Retrieval Strategies

Hasan, E.; Zhang, Y.; Cook, O.; Loe, A.; Sha, M.; T'ien, L.; Ng, M.; Rauscher, C.; Raman, S.; Bender, J. L.; Ng, R. T.; Bates, A.; Nunez, J.-J.

2026-09-04 health informatics 10.64898/2026.08.31.26361774 medRxiv
Top 1.0%
0.3%
Show abstract

Background: People affected by cancer often face difficulty finding relevant clinical, psychological, and practical support services. AI-powered navigation assistants may improve access to these resources, but their retrieval performance must be reliable. Objective: To develop a single-pass retrieval-augmented generation assistant for cancer-care navigation and compare the retrieval strategies, including their robustness to reworded questions. Methods: We created a database of 853 cancer-support resources reviewed by librarians, clinicians, researchers, and patient partners. We evaluated the system using 100 questions derived from questions submitted by patients. We compared keyword-based, semantic, and hybrid retrieval using Precision@K, Hit@K, and nDCG@K. The best-performing configuration was then tested using semantically equivalent rewordings of the original questions. Results: Keyword-based retrieval performed poorly, achieving a P@1 of 25.0% and Hit@5 of 43.0%. Semantic retrieval improved these results to 58.0% and 86.0%, respectively. The best hybrid configuration achieved a P@1 of 64.0%, Hit@5 of 90.0%, and nDCG@5 of 51.0%. Performance remained relatively stable when the questions were reworded, with a P@1 of 61.0%, Hit@5 of 88.0%, and nDCG@5 of 46.1%. Conclusions: Hybrid retrieval performed best and remained relatively stable when questions were reworded. However, its limited ability to rank a relevant resource first highlights the limitations of single-pass retrieval for patient-facing cancer navigation. Future work will explore metadata filtering and a multi-agent architecture to improve retrieval reliability.

15
The effectiveness of a complex intervention, aimed at reducing hospital occupancy, to improve Emergency Department patient flow: a retrospective controlled interrupted time series

McHenry, R. D.; Caesar, D.; Clarke, B.; Mackay, D.; Pell, J.

2026-09-03 health systems and quality improvement 10.64898/2026.08.31.26361802 medRxiv
Top 1%
0.3%
Show abstract

Objectives Emergency department (ED) crowding is recognised as an important public health concern internationally, and is driven principally by exit block, the shortage of inpatient beds for patients requiring admission. This study aimed to evaluate whether a complex intervention targeting hospital occupancy improved ED patient flow, and quantified the change in attendances. Methods A controlled interrupted time series using weekly, publicly reported Public Health Scotland data from 1 January 2022 to 1 February 2026. The multi-component intervention focused on reducing hospital occupancy and included additional adult social care funding; engagement with regional social care providers; accelerated implementation of the Discharge without Delay programme; re-evaluation of whole-hospital escalation thresholds and response; resource and data supporting inpatient department reductions in length of stay; and additional investment in remote clinical assessment. The intervention commenced at a large tertiary ED on 01 February 2025. Primary outcomes were the proportions of attendances spending [&ge;]4, [&ge;]8 and [&ge;]12 hours in the ED. The secondary outcome was attendance volume. Segmented regression was fitted with a contemporaneous control series, seasonal terms and autoregressive moving average errors. Long waits were additionally illustrated as potentially avoided deaths. Results The analysis covered 161 pre-intervention and 52 post-intervention weeks. Relative to pre-intervention levels, the proportion of attendances waiting over 4 hours fell by 10.4% (95% CI 1.6 to 19.2%), by 16.4% (95%CI 1.3 to 31.5%) over 8 hours and by 24.3% (95%CI 2.6 to 46.1%) over 12 hours. Using established associations between long ED waits and excess mortality, by one-year the intervention was potentially associated with 54 fewer excess deaths (95%CI 19 to 93). Attendances rose by 3.8% (95%CI 1.3 to 6.4%) against the counterfactual. Conclusions A complex intervention targeting hospital occupancy was associated with a reduction in long ED waits despite rising attendances. Interventions addressing hospital occupancy can meaningfully improve ED crowding.

16
Default-filled outcome labels in a deployed cognitive-screening programme: an operator-level audit and the construction of twenty-four language-model arms

Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.

2026-09-02 health informatics 10.64898/2026.08.28.26361585 medRxiv
Top 1%
0.3%
Show abstract

Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.

17
A Pragmatic Randomized Trial of an EHR-Integrated Generative AI Chart Summarization Tool for Ambulatory Clinicians

Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.

2026-08-31 health informatics 10.64898/2026.08.26.26361496 medRxiv
Top 1%
0.3%
Show abstract

BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.

18
Novel Entropy-Based Framework for Quantifying Dynamic Epistemic Uncertainty in Clinical Medicine

Yano, Y.; Shintani, E.; Arita, S.; Ashine, R.; Iinuma, N.; Mori, H.; Fujibayashi, K.; Yamada, Y.; Saita, M.; Nakashima, N.; Itoh, H.; Nangaku, M.; Ohashi, M.; Daida, H.; Arai, H.; Naito, T.

2026-08-31 health informatics 10.64898/2026.08.27.26361497 medRxiv
Top 1%
0.3%
Show abstract

The widespread adoption of clinical large language models (LLMs) introduces significant risks of automation bias, premature closure, and clinician deskilling. Current interpretability paradigms, including latent space trajectories, Concept Activation Vectors, and Concept Bottleneck Models, suffer from topological stagnation, metric distortion, and epistemic occlusion, frequently masking intermediate diagnostic uncertainty behind falsely confident outputs. To address these structural vulnerabilities, this paper introduces a novel closed-loop, multi-agent framework designed to quantify and visualize dynamic epistemic uncertainty in clinical LLM reasoning. By coupling predictive Shannon entropy with non-linear Isometric Feature Mapping (ISOMAP), the architecture projects high-dimensional inference state vectors onto a calibrated two-dimensional latent space, thereby assigning a quantifiable thermodynamic energy state to the reasoning path to track diagnostic velocity, cognitive momentum, and trajectory efficiency across sequential diagnostic rounds. Pilot validation across representative emergency medicine scenarios demonstrated distinct topological and information-theoretic behaviors: unconfounded cases (cerebellar infarction) exhibited smooth geodesic progression toward the ground truth alongside monotonic Shannon entropy decay from 2.15 to 1.74; noisy environments with ambiguous findings (spontaneous pneumothorax) suffered from trajectory wandering, local minimum traps, and high sustained entropy (~2.41) due to insufficient repulsive weighting for negative evidence; and triage-conflicted cases (acute cholangitis) achieved precise geometric proximity to the true node but experienced top-1 rank stagnation because the model conflated acute severity triage (sepsis) with anatomical etiology. By rendering machine hesitation and cognitive divergence visually auditable before final diagnostic crystallization, this geometric-information framework enables dynamic trust calibration and human-AI co-regulation at the point of care while establishing a clear mathematical foundation for future architectural interventions, such as dual-channel safety decoupling and non-linear repulsive weighting. Moving forward, validating these architectural enhancements across large-scale electronic health record databases and prospective clinical trials will be essential to realize its full clinical utility, establishing a foundational blueprint for safe, transparent, and cognitively synergistic AI integration in future medical practice. By rendering the LLM's reasoning process visually auditable, this framework lays the groundwork for capturing and externalizing the clinician's own cognitive patterns within the AI, forming a coupled system. This enables the explicit visualization of cognitive gaps between physician hypotheses and AI inferences, transforming the interaction from simple answer-checking into a dynamic learning process for both human and machine that prevents diagnostic oversight. Ultimately, because the responsibility for final clinical decision-making remains with the human practitioner, this framework serves as a vital decision-support mechanism. Moving forward, validating these architectural enhancements across large-scale electronic health record databases and prospective clinical trials will be essential to realize its full clinical utility, establishing a foundational blueprint for safe, transparent, and cognitively synergistic AI integration in future medical practice.

19
Addressing Measurement Error of Machine-Learned Physical Activity in Nonlinear Dose-Response Survival Analysis: Development and Evaluation of Accelerated Failure Time, Spline, and Simulation-Extrapolation Method

Mamiya, H.; Zhang, Q.; Zhang, X.; Yan, Y.; Sharma, A.

2026-08-31 epidemiology 10.64898/2026.08.25.26361155 medRxiv
Top 1%
0.2%
Show abstract

Wearable (accelerometer) data and machine-learning allow objective assessment of the amount of daily physical activity. However, wearable-derived human activity is subject to measurement error. No studies have corrected the dose-response association between physical activity and survival time to chronic diseases, including cardiovascular disease (CVD). The objective is to estimate the measurement error-corrected association between CVD events and multiple measures of daily duration of light and total physical activity, derived from machine-learning and conventional accelerometer-processing methods. Our method combined an accelerated failure time model, spline, and simulation-extrapolation (SIMEX). The method recovered the true dose-response non-linear association in simulated data, while the naive model failed to capture it due to substantial attenuation. Application to the UK Biobank accelerometer cohort also showed an increased protective association of total physical activity after SIMEX correction (Time Ratio [TR] = 1.56, 95% CI: 1.28-1.82 vs. TR = 1.38, 95% CI: 1.24-1.54 for SIMEX-corrected vs. uncorrected dose-response association between the 95th and 5th percentiles of total activity), with a similar increase for light physical activity. Sensitivity analysis indicates that the female population experiences a substantially larger protective association after SIMEX correction than males. Dose-response survival analysis is a widely used analytical method in physical activity epidemiology and benefits from measurement error correction.

20
Skin Cancer Classification Using Explainable Artificial Intelligence With an Ensemble Model and Rigorous Leakage Free Validation

BARAN, M. T.; KARAKOYUN, O.

2026-09-05 health systems and quality improvement 10.64898/2026.09.02.26362011 medRxiv
Top 1%
0.2%
Show abstract

Background: Reliable melanoma classification requires models that capture both local dermoscopic morphology and broader contextual patterns while maintaining auditable, leakageaware internal validation. Objectives: To develop and internally validate an EfficientNetB0-Swin Transformer Tiny ensemble for classifying histopathologically verified dermoscopic images as benign melanocytic lesions or malignant melanoma. Methods: This retrospective diagnostic model-development and internal validation study screened 552,869 ISIC Archive records; filtering and dermatologist review yielded 1,199 uniquepatient and unique lesion images (578 benign and 621 malignant). Images were the predictors and histopathology was the reference. ImageNet pretrained EfficientNetB0 and Swin-T features were fused. Patient independent five fold validation used weighted sampling, mixup, label smoothing, AdamW, early stopping, and five view test time augmentation. Results: Mean accuracy was 0.89325 {+/-} 0.03179, mean receiver operating characteristic area under the curve (ROC-AUC) was 0.96348 {+/-} 0.01695, and mean support weighted F1-score was 0.89300 {+/-} 0.03220. The fold level 95% confidence intervals were 0.8538-0.9327 for accuracy and 0.9424-0.9845 for ROC-AUC. Pooled counts were 526 true negatives, 52 false positives, 76 false negatives, and 545 true positives, yielding 87.76% sensitivity and 91.00% specificity. Qualitative Grad-CAM review showed peripheral artifact activation in two false positives and lesion centered activation in two correctly classified cases; these observations were not systematically scored. Limitations: The validation folds were also used for early stopping and checkpoint selection. Device stratified analysis, systematic interpretability scoring, calibration, and independent external validation were unavailable. Conclusions: The ensemble showed high internal discrimination and is intended only as a clinician facing adjunct. The error audit workflow enables targeted retrospective review, but external validation is required before clinical use or generalizability claims.